Papers with cultural sensitivity

9 papers
Tailored Emotional LLM-Supporter: Enhancing Cultural Sensitivity (2026.eacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown growing potential in offering emotional support, but their ability to deliver culturally sensitive support remains underexplored due to a lack of resources.
Approach: They propose a large language model dataset that includes 1,729 distress messages, 1,523 cultural signals and 1,041 support strategies with fine-grained emotional and cultural annotations.
Outcome: The proposed models outperform peer-reviewed models and lack cultural sensitivity.
RENOVI: A Benchmark Towards Remediating Norm Violations in Socio-Cultural Conversations (2024.findings-naacl)

Copied to clipboard

Challenge: Norm violations occur when individuals fail to conform to culturally accepted behaviors, which may lead to potential conflicts.
Approach: They propose to use a large corpus of 9,258 multi-turn dialogues annotated with social norms to equip AI systems with a remediation ability.
Outcome: The proposed system can understand and remediate norm violations step by step.
EtiCor++: Towards Understanding Etiquettical Bias in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Etiquettes are region-specific and are an essential part of the culture of a region.
Approach: They introduce EtiCor++, a corpus of etiquettes worldwide, to evaluate LLMs for their knowledge about etiques across regions.
Outcome: The proposed corpus of etiquettes shows that LLMs are biased towards certain regions.
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore’s Low-Resource Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have transformed natural language processing, but their safety mechanisms remain under-explored in low-resource, multilingual settings.
Approach: They propose a red-teaming approach to probe LLM vulnerabilities in Singapore's diverse linguistic context using a dataset and evaluation framework.
Outcome: The proposed framework systematically probes LLM vulnerabilities in three real-world scenarios including Singlish, Chinese, Malay, and Tamil.
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition (2025.findings-acl)

Copied to clipboard

Challenge: Existing research has evaluated large language models' cultural knowledge and contextual understanding, reducing their effectiveness in multicultural settings.
Approach: They propose a benchmark to evaluate LLMs' cultural understanding with a focus on Korean superstitions.
Outcome: The proposed benchmark assesses multilingual LLMs in Korean and English to analyze their ability to reason about Korean cultural contexts and how language variations affect performance.
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation datasets lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage.
Approach: They propose to use multilingual consistency as a complementary metric to assess performance bottlenecks and guide model improvement.
Outcome: The proposed model lacks cross-lingual alignment and language coverage gaps between state-of-the-art models.
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a benchmark is designed to assess the comprehension of Arabic poetry by large language models in 12 historical eras.
Approach: They propose a benchmark to assess the comprehension of Arabic poetry by large language models in 12 historical eras.
Outcome: The benchmark assesses the comprehension of Arabic poetry by large language models in 12 historical eras.
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users (2025.acl-long)

Copied to clipboard

Challenge: Despite high adoption rate of Large Language Models, there are limitations related to contextual understanding, cultural sensitivity, and complex scene understanding.
Approach: They conduct a user survey to identify adoption patterns and key challenges users face with such technologies.
Outcome: The proposed models have high adoption rates but still face limitations in visual aids.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations